Ops efficiency followup result
2026-09-28 11:45:19 EDT · rdmsm4x · ops@rdmsm4x/opseff0928
Work began 2026-09-28 11:25:22 EDT. TASK-20260928-08; child
TASK-20260928-09; rollout decision DEC-20260928-03. Rich explicitly
authorized integration, source fixes, controller non-objection, tests,
rollback and dry-run measurement. No human messages, purchases,
credential changes, irreversible deletion, or session termination. Usage
gate returned conserve; no workers spawned. Heavy
verification ran at nice 10.
Delivered and integrated
- TASK-20260928-07: census tool and nine fixture
tests integrated into
/Users/richh/dev/fleet/maintenance/tools/efficiency/. - DOC-20260928-02 / DOC-20260928-03: prior-fix and claim/stall worksheets integrated beside the tool. Historical evidence links corrected for the canonical location.
- Maintenance implementation commits: e920114, 5da5612.
- Ticket implementation commits: a5c224f, a78dc8c; merged into canonical issues main via b1a2327, 04b9fbd. Existing dirty files and automatic ticket commits preserved.
- Branches integrated into canonical main; both
fleetandbackupremotes pushed non-force. Final record commits and remote verification are recorded below.
Status producers and deployment gap
The 1,785 baseline messages are seven known machine-status classes; their bodies were not all identical. TASK-20260927-72 had already migrated canonical producers to telemetry. This followup found all five installed spoke coordinators still using inbox mail. Only their bus-call argument list was backported; unrelated v1.15 behavior was preserved.
| Class | Baseline | Canonical producer | Result |
|---|---|---|---|
| fleetstate-* | 926 | ~/scripts/fleet_state_report.zsh |
Already telemetry on six hosts; verified |
| agent-coordinator-* | 428 | fleet/agent-coordinator/agent_coordinator.py; installed
under ~/.agent-coordination/coordinator/bin/ |
Hub telemetry verified; five spokes repaired |
| devmon-daemon-restart* | 125 | dev/_ops/devmon/devmon.sh |
Already telemetry on six hosts; verified |
| agent-memory-advisory* | 123 | same devmon producer | Already telemetry on six hosts; verified |
| uncommitted-work-in-* | 106 | fleet/maintenance/scripts/shared_tree_guard.zsh |
Hub telemetry producer verified |
| fleet-audit* | 59 | dev/scripts/fleet_audit.zsh |
Hub telemetry producer verified |
| devmon-daemon-alert* | 18 | same devmon producer | Already telemetry on six hosts; verified |
All six installed buses now have exact-body append dedupe. Latest snapshots retain a fresh timestamp on every observation; changed values and recoveries still append. No status/PID/number normalization hides changed alerts. Existing telemetry filenames, frontmatter and replication remain intact; real agent-to-agent mail is unchanged. Additional machine-topic scan found actionable ticket-store push/security failures and CVE alerts, which were retained; stability-review already uses telemetry.
| Dry-run measurement | Before | After |
|---|---|---|
| Historical seven-class status messages entering inbox | 1,785 | 0 |
| Latest telemetry snapshots from that replay | — | 36 |
| Historical append records | 1,785 | 1,785 |
| Identical-body fixture append records | 2 | 1 |
| Installed spoke coordinators still sending mail | 5 | 0 |
The replay used actual agent_msg.zsh in an isolated
HOME/coordination root with sync disabled. All 1,785 historical bodies
differed, so no historical log-volume reduction is claimed from
exact-body dedupe. The mailbox reduction includes the earlier
migration; this task completed its missing deployment. This is not an
observed future-week saving. Six remote hash checks and six isolated
producer-call smoke checks passed. We did not run full coordinator jobs,
restart agents, or claim a naturally scheduled pass as proof.
Evidence: inventory before, inventory after, replay, rollout, spoke rollback, fleet verification. Source replay script is in the canonical kit.
Abandoned claim causes and source fixes
The baseline has 81 force-release events; 73
explicitly describe stale/dead/expired claims. This is 70 distinct
tickets, not 73 active sessions. Among those 73, former holders were
Claude 37, Codex 5, Tyrell 4, xattr 3, launch 3,
domains 3, rdmactccfix 2, library 2, and 14 other single-event lanes.
Largest individual holders across all 81:
claude@rdmsm4x/tyrlead1 12 and
claude@rdmsm4x/L-ops 8.
The top recorded causes are session exit/lapse with no heartbeat or closeout: 36 events in the Sep 27 stale/dead batch and 13 in the Sep 28 lapsed/no-heartbeat batch. Four more cite >12h without a heartbeat; four XEntropy releases cite post-reboot process/worktree evidence. Smaller cases name completed headless workers, a failed driver, blocked owner operations, and handoff/closeout omissions. Bot lane names do not establish their underlying harness. Historical release reasons cannot distinguish every true exit from false staleness. See claim analysis for per-event identity/reason evidence.
Two source defects were addressed:
- Renewal lost on reindex.
_extendupdated only derived SQLite. A comment/reindex restored the acquisition expiry from Markdown. Three regression cases failed before; all now pass. Heartbeat snapshots persist/coalesce; original acquisition and native provenance survive rebuilds. Failed persistence fails renewal instead of reporting success. - No reliable exit closeout. Hub Claude had no
SessionEnd claim cleanup, and named ticket sessions hid native
transcript provenance. Acquisition now records an exact, unambiguous
native filename through metadata-only lookup; no native contents are
read. The installed SessionEnd hook releases only that exact transcript
on that exact host.
--if-transcriptrechecks provenance under the ticket write lock. No--force, no Stop cleanup, no parent/subagent cleanup, no task resolution, no process control.
| Claim verification | Before | After |
|---|---|---|
| Heartbeat persistence regressions | 3 failing / 5 cases | 0 failing; expanded to 8 passing cases |
| Isolated terminating-session claim | 1 | 0 |
| Same claim after ordinary Stop | 1 | 1 |
| Task status after automatic release | in-progress | in-progress |
The cleanup targets the largest measured harness group, hub only. New Claude sessions load the installed hook; existing sessions were not restarted. Natural session-end capture has not yet been observed. Missing/ambiguous provenance is deliberately skipped; crashes, SIGKILL and power loss still fall back to bounded leases. No claim of 73 prevented events or measured next-week improvement is made.
Verification totals and rollback
232 passing automated checks: 130 bus, 69 existing claim/lease, 8 new renewal/provenance, 10 cleanup/installer, 6 telemetry, 9 census. Also: six host hash checks, six isolated producer smokes, and the installed-hook isolated lifecycle scenario. Original bus suite exposed a pre-existing test leak into the live localdb reader and assumed host ordering; fixture isolation and explicit host selection fixed it. Final suite is 130/130, not a waived failure.
- bus suite
- claim suite
- renewal regressions
- cleanup tests and installed lifecycle
- telemetry tests and census tests
One-step fleet telemetry and hub-hook rollback (preserves backups and refuses later drift):
zsh /Users/richh/dev/fleet/maintenance/scripts/rollback_efficiency.zsh --applyIndividual host rollback:
python3 ~/dev/fleet/maintenance/scripts/deploy_efficiency.py --host HOST --rollback.
Hub hook only:
python3 ~/dev/fleet/maintenance/scripts/install_session_claim_cleanup.py --rollback.
Hub and spoke backup/rollback/reapply paths were exercised. Backups on
each host: ~/.agent-coordination/backups/opseff0928/ with
before/after SHA-256 manifest. The wrapper itself was syntax checked and
dry-run verified; full simultaneous rollback was not run. Independent
one-step claim rollback:
zsh ~/dev/fleet/maintenance/scripts/rollback_claim_efficiency.zsh --apply.
It removes only this hook and creates non-force Git reverts of a78dc8c
and a5c224f, preserving ticket data and history. The wrapper refuses
edited source paths, was syntax/dry-run checked, and has not been
executed on canonical issues. Commits remain available for
reapplication.
Limits and durable closeout
The requested old topology-audit path is absent; its canonical
replacement at
fleet/dev-fleet-reconciliation/scripts/audit_codex_fleet_topology_v1.0.zsh
ran. Six hosts were reachable; three topology expectations failed
(rdmpw3265m account-lane mismatch, jdmbair13m5 unclassified account lane
and project-root state). No account/config/history was moved or altered
to address those adjacent findings. Audit.
DEC-20260928-03 was sent to the controller before mutation; no
objection received. Its rollout record was appended to
SETTLED-ANSWERS.md. Existing unrelated dirty work retained. Apple Notes
publication is pending:
launchctl managername reports Background. The
fleet-notes-publish rule requires a file copy in this case; it is saved
in the changelog archive for the existing Notes autopublisher. No human
intervention requested.
Closeout update: source race guard merged at maintenance 5da5612 and issues 04b9fbd; TASK-20260928-09 and DEC-20260928-03 resolved with evidence. Artifact tickets carry canonical integration receipts. Evidence/rollback record commit: 9fd3b29. Ticket closeout commit: 96e540a. Both main tips were read back from fleet and backup and matched exactly at closeout.
TASK-20260928-08 is resolved and both task leases are released. Notes tool returned exit3 (Background, not Aqua); file archive: /Users/richh/dev/LLM/Claude/changelogs/rdmsm4x - ops - Efficiency source fixes - codex - TASK2026092808 - fleet - 20260928-1147.md.
Final verification: zero claims for ops@rdmsm4x/opseff0928. Worktrees
preserved under
~/dev/_worktrees/archive/ops-eff0928-maintenance and
ops-eff0928-issues; no files deleted. Installed guarded
SessionEnd lifecycle was rerun successfully after the race-guard merge.
Internal completion receipt: bus 20260928-114839-8CCE30B8.